Papers with task completion rate

8 papers
Bootstrapping a Neural Conversational Agent with Dialogue Self-Play, Crowdsourcing and On-Line Reinforcement Learning (N18-3)

Copied to clipboard

Challenge: End-to-end neural models for conversational agents require large corpus of dialogues to learn effectively.
Approach: They propose a method for building an agent for arbitrary tasks by combining dialogue self-play and crowd-sourcing.
Outcome: The proposed approach can be quickly bootstrapped to deploy in front of users and further optimized via interactive learning from actual users.
Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation (2021.eacl-main)

Copied to clipboard

Challenge: Outdoor vision-and-language navigation (VLN) tasks require visual grounding to generate correct actions.
Approach: They propose a multimodal text style transfer learning approach to mitigate data scarcity in outdoor vision-and-language navigation tasks.
Outcome: The proposed approach outperforms baseline models on the outdoor vision-and-language navigation task, improving task completion rate by 8.7% relative to the baseline models.
LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferences (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study shows that large language models (LLMs) are limited in understanding natural language preferences.
Approach: They propose a novel LLM-as-Parser-based route planning system that utilizes an LLM to comprehend natural language, extract user preferences and recognize task dependencies.
Outcome: The proposed system achieves superior performance with guarantees across multiple constraints.
VoxMind: An End-to-End Agentic Spoken Dialogue System (2026.acl-long)

Copied to clipboard

Challenge: Existing research on end-to-end spoken dialogue models has focused on core perception and generation, with limited exploration of tool-augmented extensions.
Approach: They propose a framework to equip end-to-end spoken dialogue models with comprehensive agentic abilities by leveraging a 470-hour AgentChat dataset.
Outcome: The proposed framework outperforms Gemini-2.5-Pro on spoken agent tasks while maintaining general conversational quality.
InferPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents (2026.findings-acl)

Copied to clipboard

Challenge: Inference attacks are important for assessing model's robustness, but their implementation and parameters are challenging for non-experts.
Approach: They propose an autonomous agent capable of conducting inference attacks without human intervention.
Outcome: The proposed agent achieves a 100.0% task completion rate and near-expert attack performance with an average token cost of only 0.627 per run.
MLAlgo-Bench: Can Machines Implement Machine Learning Algorithms? (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, the top-performing models achieve a 48.8% task completion rate on realizing machine learning algorithms .
Approach: They propose a benchmark to test machine learning's ability to generate ML code for humans . they propose an automatic evaluation framework with metrics such as task pass rate and time overhead .
Outcome: The proposed benchmark is unique in its focus on interpreting complex human instructions and producing multi-step, high-complexity code.
Retrospective Learning from Interactions (2025.acl-long)

Copied to clipboard

Challenge: Multi-turn interactions between large language models and users naturally include implicit feedback signals.
Approach: They propose a method to learn from feedback signals in past interactions without annotations . they use a multimodal LLM to solve a reasoning task with a combinatorial solution space .
Outcome: The proposed method improves task completion rate from 31% to 82% without annotations.
TPS-Bench: Evaluating AI Agents’ Tool Planning & Scheduling Abilities in Compounding Tasks (2026.acl-long)

Copied to clipboard

Challenge: Large language model (LLM) agents have demonstrated strong problem-solving competence across domains like research and coding.
Approach: They propose to use a tool repository to analyze the ability of large language model agents to solve complex problems.
Outcome: The proposed model outperforms open-source and closed-source models in task completion rate and efficiency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations